Papers with Graphical User Interfaces
SLATE: A Super-Lightweight Annotation Tool for Experts (P19-3)
Copied to clipboard
| Challenge: | a new annotation tool is designed to fill the niche of a lightweight interface for terminal users . current tools are built with direct manipulation via a Graphical User Interface (GUI) this approach is time-consuming and difficult to modify . |
| Approach: | They propose a terminal-based annotation tool that supports multiple annotations . they use a text-based interface that uses almost the entire screen to display documents . |
| Outcome: | The proposed tool is designed to fill the niche of a lightweight interface for users with a terminal-based workflow. |
VGA: Vision GUI Assistant - Minimizing Hallucinations through Image-Centric Fine-Tuning (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing Large Vision-Language Models (VLMs) often overly rely on internal text-based knowledge while neglecting visual inputs. |
| Approach: | They propose a model that balances attention image and text to enhance interpretation and reduce hallucinations by using a visual input. |
| Outcome: | The proposed model improves interpretation and reduces hallucinations by balancing attention image and text to enhance interpretation and reduction of hallucinosity. |
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding (2025.findings-acl)
Copied to clipboard
Joonhyung Park, Peng Tang, Sagnik Das, Srikar Appalaraju, Kunwar Yashraj Singh, R. Manmatha, Shabnam Ghadar
| Challenge: | Existing vision-only GUI agents ground elements from large and cluttered screenshots, requiring them to process substantial irrelevant information that compromises their accuracy. |
| Approach: | They propose a visual agent model for GUI automation that leverages zoomed-in region proposals for precise element localization. |
| Outcome: | The proposed approach improves state-of-the-art grounding accuracy by 13% across diverse GUI platforms on the GUI grounding benchmarks ScreenSpot and AgentStudio. |
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for GUI understanding ignore a key GUI-referring task: screen reading based on user-indicated points. |
| Approach: | They propose a Tree-of-Lens agent that constructs a Hierarchical Layout Tree based on user input points and a GUI screenshot. |
| Outcome: | The proposed agent can interpret the Screen Point-and-Read task on mobile, web, and operating systems. |
LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization (2026.findings-acl)
Copied to clipboard
Jiaqi Tang, Yu Xia, Yi-Feng Wu, Yuwei Hu, Chen Yuhui, Qing-Guo Chen, Xiaogang Xu, Xiangyu Wu, Hao LU, Yanqing Ma, Shiyin Lu, Qifeng Chen
| Challenge: | Existing strategies for spatial localization are limited due to their limited capacity to perceive positional data. |
| Approach: | They propose a location-based approach that leverages locational data to optimize interaction preferences. |
| Outcome: | The proposed approach achieves SOTA results across offline benchmarks and real-world evaluations. |